Papers with Visual QA
Being Negative but Constructively: Lessons Learnt from Creating Better Visual Question Answering Datasets (N18-1)
Copied to clipboard
| Challenge: | Visual question answering datasets are a form of (visual) Turing test that artificial intelligence should strive to achieve. |
| Approach: | They propose automatic procedures to remedy design deficiencies in visual question answering datasets . they propose to use a set of decoys to re-construct decoying answers for two popular Visual QA datasets. |
| Outcome: | The proposed procedures improve the performance of the proposed datasets. |
RoD-TAL: A Benchmark for Answering Questions in Romanian Driving License Exams (2026.findings-eacl)
Copied to clipboard
Andrei Vlad Man, Răzvan-Alexandru Smădu, Cristian-George Craciun, Dumitru-Clementin Cercel, Florin Pop, Mihaela-Claudia Cercel
| Challenge: | a growing need for tools that support legal education, especially in under-resourced languages such as Romanian . we evaluate the capabilities of large language models and vision-language models in legal education . |
| Approach: | They evaluate the capabilities of Large Language Models and Vision-Language Models in Romanian driving law through textual and visual question-answering tasks. |
| Outcome: | The proposed model improves retrieval performance and QA accuracy in Romanian driving tests. |
***YesBut***: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language Models (2024.emnlp-main)
Copied to clipboard
Abhilash Nandy, Yash Agarwal, Ashish Patwa, Millon Das, Aman Bansal, Ankit Raj, Pawan Goyal, Niloy Ganguly
| Challenge: | Existing Vision-Language models perform poorly on satirical image detecting tasks . satire and humor are powerful tools to highlight issues, provoke thought, and encourage critical perspective . |
| Approach: | They propose to use a dataset to evaluate satirical images and satire images to detect satiric images . they also propose to generate the reason behind the image being satiral by generating one half of the image to be satisfying . |
| Outcome: | The proposed dataset contains 2547 images, 1084 satirical and 1463 non-satirically, with different artistic styles. |